Papers with conversational texts
Designing the Business Conversation Corpus (D19-52)
Copied to clipboard
| Challenge: | Existing parallel corpora for machine translation of written text and monologues are limited. |
| Approach: | They propose to introduce a Japanese-English business conversation parallel corpus into machine translation training scenarios and show how it improves machine translation quality. |
| Outcome: | The proposed corpus is used in a Japanese-English business conversation training scenario and shows how it performs. |
SuperDialseg: A Large-scale Dataset for Supervised Dialogue Segmentation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Empirical studies show that supervised learning is extremely effective in in-domain datasets and models trained on SuperDialseg can achieve good generalization ability on out-of-domain data. |
| Approach: | They propose a supervised definition of dialogue segmentation points using document-grounded dialogues and a large-scale supervised dataset called SuperDialseg. |
| Outcome: | The proposed model can achieve good generalization ability on out-of-domain data. |
Introducing a Large-Scale Dataset for Vietnamese POS Tagging on Conversational Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | POS taggers are trained on informal texts which contain many informal inputs such as acronyms, abbreviations, out-of-vocabulary words, etc. |
| Approach: | They propose a large-scale human-labeled dataset for the Vietnamese POS tagging task on conversational texts and develop an annotation guideline to manually annotate 16.310K sentences using this guideline. |
| Outcome: | The proposed tagging scheme achieved 93.36% accuracy score and higher than the model with handcrafted features and fine-tuning BERT. |
Discriminating between Similar Languages on Imbalanced Conversational Texts (L18-1)
Copied to clipboard
| Challenge: | Empirical results suggest that our system achieves an accuracy of 95.7% on our Uyghur and Kazakh dataset, which is higher than that of the CNN classifier. |
| Approach: | They propose to build a balanced Uyghur and Kazakh corpus and build morphological classifiers to discriminate between the two languages. |
| Outcome: | The proposed system outperforms the champions on both test sets B1 and B2. |
Linear Semantic Segmentation for Low-Resource Spoken Dialects (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing models for semantic segmentation are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resourced conversational varieties. |
| Approach: | They propose a multi-genre benchmark for semantic segmentation in Arabic, focusing on dialectal discourse. |
| Outcome: | The proposed model outperforms baselines on dialectal non-news genres while performing well on high-resource written text. |